iT邦幫忙

2026 iThome 鐵人賽

DAY 16
0

前言

Day 15 我們把 Agent 從 FakeLLMClient 接到 Gemini Flash。

接上真正的 LLM 後,evaluation result 開始變得更接近真實情境。

例如原本 fake client 可能只會回:

Fake response for: 請回答 HTTP 狀態碼 404 通常代表什麼

接上 Gemini 後,模型可能會正確回答:

HTTP 狀態碼 404 通常代表找不到請求的資源。

這讓 success rate 從 fake baseline 明顯提升。

但 Day 15 的結果也留下了一些很有價值的失敗案例。

例如 JSON 題的實際輸出可能是:

```json
{
  "answer": 15
}
```

人類看得出來它其實是正確 JSON。

但對 json.loads() 來說,整段文字不是合法 JSON,因為外面多了 markdown code fence。

所以 evaluator 會判定:

Output is not valid JSON: Expecting value

到目前為止,平台只知道:

passed: false
failure_reason: Output is not valid JSON: Expecting value

這裡還可以再補一層。

我們要讓平台不只記錄失敗原因,也要記錄失敗類型。


今天要完成什麼?

這一篇要做到的是:

在 evaluation result 中加入 failure_type。

會完成幾件事:

  1. 定義本系列會使用的 Agent failure types。
  2. 修改 evals/evaluators.py 的 EvaluationResult。
  3. 實作 rule-based failure type classifier。
  4. 修改 evaluator,讓失敗結果包含 failure_type。
  5. 修改 evals/runner.py,讓輸出的 JSON 包含 failure_type。
  6. 重新執行 Gemini baseline,觀察失敗分布。

先不碰這些:

  • 自動修復失敗。
  • retry。
  • dashboard 圖表。
  • LLM-as-a-Judge。
  • 完整 evaluator false positive 偵測。

範圍先收斂在一件事:

把失敗從一段文字,整理成可以統計的結構化類別。


為什麼需要 Failure Type?

failure_reason 適合給人看。

例如:

Expected output to contain '任務', but got 'AI Agent 是一種能夠自主感知環境...'

這句話很清楚,但不適合做統計。

如果我們想回答:

這次 eval run 裡,最多的是格式錯誤,還是答案錯誤?

只靠 failure_reason 會很麻煩。

因為每一筆 reason 都可能長得不一樣。

例如:

Expected exactly 'OK', but got 'OK。'
Output is not valid JSON: Expecting value
Missing required key: answer
Unknown tool: search
Agent execution error

這些字串適合除錯,但不適合分析。

所以我們需要一個更穩定的欄位:

{
  "passed": false,
  "failure_type": "format_error",
  "failure_reason": "Output is not valid JSON: Expecting value"
}

failure_type 的用途是讓平台可以統計。

failure_reason 的用途是讓人可以追原因。

兩者不是互相取代,而是互補。


本系列的 Failure Types

先定義六種失敗類型。

failure_type 說明
wrong_answer Agent 有回答,但答案不符合預期
format_error 輸出格式錯誤,例如不是合法 JSON
instruction_error 沒有遵守明確指令,例如要求只回覆 OK
tool_error 工具呼叫失敗、工具不存在或工具輸入錯誤
execution_error Agent 執行過程發生 exception、timeout 或 API error
unknown_failure 暫時無法分類的錯誤

另外還有一種情況要特別說明:

evaluator_false_positive

這表示 evaluator 判定通過,但其實不該通過。

例如 Day 14 提過的案例:

問題:請回答 Python 中 list 是可變還是不可變資料型別
expected:可變
actual:Fake response for: 請回答 Python 中 list 是可變還是不可變資料型別

因為 actual 裡剛好包含「可變」,所以 contains 判定通過。

但這不是 Agent 真正答對,而是 evaluator 誤判。

這種情況比較難靠單純 rule-based 自動偵測。

所以今天先把它放在概念分類中,後面做 dashboard 和人工分析時再討論。

這次的實作重點先放在:

failed case -> failure_type

今天的專案結構

這次會修改兩個檔案。

agent-testing-platform/
  evals/
    __init__.py
    cases.json
    runner.py
    evaluators.py

會修改:

檔案 修改內容
evals/evaluators.py 讓 EvaluationResult 包含 failure_type,並加入分類邏輯
evals/runner.py 把 failure_type 寫入每一筆 eval result

不修改:

  • evals/cases.json
  • agents/simple_agent.py
  • agents/gemini_llm.py
  • agents/client_factory.py

因為這一篇要分類的是 evaluation result,不是改變 Agent 行為。


修改 EvaluationResult

Day 12 建立的 EvaluationResult 目前長這樣:

@dataclass
class EvaluationResult:
    passed: bool
    failure_reason: str | None = None

這次要多加一個欄位:

failure_type

修改 evals/evaluators.py:

@dataclass
class EvaluationResult:
    passed: bool
    failure_reason: str | None = None
    failure_type: str | None = None

failure_type 可以是 None。

通過的案例不需要失敗類型。

例如:

{
  "passed": true,
  "failure_type": null,
  "failure_reason": null
}

而失敗的案例會有明確分類:

{
  "passed": false,
  "failure_type": "format_error",
  "failure_reason": "Output is not valid JSON: Expecting value"
}

實作 Failure Type Classifier

接著實作一個簡單的分類 function。

這個 function 會接收:

  • test_case
  • failure_reason

然後回傳對應的 failure_type。

修改 evals/evaluators.py,新增 classify_failure():

def classify_failure(test_case: dict, failure_reason: str | None) -> str:
    grading_method = test_case["grading_method"]
    task_type = test_case["task_type"]
    reason = failure_reason or ""

    if "not valid JSON" in reason:
        return "format_error"

    if "Missing required key" in reason:
        return "format_error"

    if "JSON output must be an object" in reason:
        return "format_error"

    if grading_method == "json_exact":
        return "format_error"

    if task_type == "instruction_following":
        return "instruction_error"

    if "Unknown tool" in reason:
        return "tool_error"

    if "Unsupported grading method" in reason:
        return "unknown_failure"

    if task_type in {"keyword_qa", "general_qa", "calculation"}:
        return "wrong_answer"

    return "unknown_failure"

這個 classifier 先採 rule-based。

它不是完美分類器,但很適合 MVP 階段。

例如只要 failure_reason 包含:

not valid JSON

就分類成:

format_error

如果任務類型是:

instruction_following

而且沒有通過,就分類成:

instruction_error

如果是一般問答、關鍵字問答或計算題沒有通過,先分類成:

wrong_answer

要注意的是:wrong_answer 不一定代表模型真的完全錯。

例如 Day 15 的 case_008:

請用一句話說明什麼是 AI Agent

Gemini 的回答語意上可能合理,但因為沒有包含 expected 裡指定的「任務」兩個字,所以被判定失敗。

這時候 wrong_answer 比較精確地說,是:

在目前 evaluator 規則下,答案不符合預期。

後面我們可以再細分成:

  • true wrong answer
  • evaluator too strict
  • semantic equivalent but keyword missing

但這一篇先不要把分類做太細。


讓 evaluator 回傳 failure_type

現在要讓每個 evaluator 在失敗時都帶上 failure_type。

一種做法是在每個 evaluator 裡自己判斷。

例如:

return EvaluationResult(
    passed=False,
    failure_reason="...",
    failure_type="format_error",
)

但這樣會讓 evaluate_exact_match()、evaluate_contains()、evaluate_json_exact() 裡到處重複分類邏輯。

所以這裡採用另一種做法:

  1. 各 evaluator 只負責產生 passed 和 failure_reason。
  2. 統一在 evaluate() 最後補上 failure_type。

修改 evals/evaluators.py 的 evaluate():

def evaluate(test_case: dict, actual: str | None) -> EvaluationResult:
    expected = test_case["expected"]
    grading_method = test_case["grading_method"]

    if grading_method == "exact_match":
        result = evaluate_exact_match(expected, actual)
    elif grading_method == "contains":
        result = evaluate_contains(expected, actual)
    elif grading_method == "json_exact":
        result = evaluate_json_exact(expected, actual)
    else:
        result = EvaluationResult(
            passed=False,
            failure_reason=f"Unsupported grading method: {grading_method}",
        )

    if not result.passed:
        result.failure_type = classify_failure(
            test_case=test_case,
            failure_reason=result.failure_reason,
        )

    return result

這樣設計可以把分類邏輯集中在一個地方。

未來要調整 failure type 規則時,不需要去每個 evaluator 裡面找。


evals/evaluators.py 完整版本

整理後,evals/evaluators.py 會像這樣。

修改 evals/evaluators.py:

import json
from dataclasses import dataclass
from typing import Any


@dataclass
class EvaluationResult:
    passed: bool
    failure_reason: str | None = None
    failure_type: str | None = None


def evaluate_exact_match(expected: Any, actual: str | None) -> EvaluationResult:
    if actual is None:
        return EvaluationResult(
            passed=False,
            failure_reason="Actual output is None",
        )

    expected_text = str(expected).strip()
    actual_text = actual.strip()

    if actual_text == expected_text:
        return EvaluationResult(passed=True)

    return EvaluationResult(
        passed=False,
        failure_reason=f"Expected exactly '{expected_text}', but got '{actual_text}'",
    )


def evaluate_contains(expected: Any, actual: str | None) -> EvaluationResult:
    if actual is None:
        return EvaluationResult(
            passed=False,
            failure_reason="Actual output is None",
        )

    expected_text = str(expected).strip()

    if expected_text in actual:
        return EvaluationResult(passed=True)

    return EvaluationResult(
        passed=False,
        failure_reason=f"Expected output to contain '{expected_text}', but got '{actual}'",
    )


def parse_json_output(actual: str | None) -> tuple[dict[str, Any] | None, str | None]:
    if actual is None:
        return None, "Actual output is None"

    try:
        parsed = json.loads(actual)
    except json.JSONDecodeError as exc:
        return None, f"Output is not valid JSON: {exc.msg}"

    if not isinstance(parsed, dict):
        return None, "JSON output must be an object"

    return parsed, None


def evaluate_json_exact(expected: Any, actual: str | None) -> EvaluationResult:
    if not isinstance(expected, dict):
        return EvaluationResult(
            passed=False,
            failure_reason="Expected value for json_exact must be an object",
        )

    parsed, error = parse_json_output(actual)

    if error:
        return EvaluationResult(
            passed=False,
            failure_reason=error,
        )

    for key, expected_value in expected.items():
        if key not in parsed:
            return EvaluationResult(
                passed=False,
                failure_reason=f"Missing required key: {key}",
            )

        actual_value = parsed[key]

        if actual_value != expected_value:
            return EvaluationResult(
                passed=False,
                failure_reason=(
                    f"Expected key '{key}' to be '{expected_value}', "
                    f"but got '{actual_value}'"
                ),
            )

    return EvaluationResult(passed=True)


def classify_failure(test_case: dict, failure_reason: str | None) -> str:
    grading_method = test_case["grading_method"]
    task_type = test_case["task_type"]
    reason = failure_reason or ""

    if "not valid JSON" in reason:
        return "format_error"

    if "Missing required key" in reason:
        return "format_error"

    if "JSON output must be an object" in reason:
        return "format_error"

    if grading_method == "json_exact":
        return "format_error"

    if task_type == "instruction_following":
        return "instruction_error"

    if "Unknown tool" in reason:
        return "tool_error"

    if "Unsupported grading method" in reason:
        return "unknown_failure"

    if task_type in {"keyword_qa", "general_qa", "calculation"}:
        return "wrong_answer"

    return "unknown_failure"


def evaluate(test_case: dict, actual: str | None) -> EvaluationResult:
    expected = test_case["expected"]
    grading_method = test_case["grading_method"]

    if grading_method == "exact_match":
        result = evaluate_exact_match(expected, actual)
    elif grading_method == "contains":
        result = evaluate_contains(expected, actual)
    elif grading_method == "json_exact":
        result = evaluate_json_exact(expected, actual)
    else:
        result = EvaluationResult(
            passed=False,
            failure_reason=f"Unsupported grading method: {grading_method}",
        )

    if not result.passed:
        result.failure_type = classify_failure(
            test_case=test_case,
            failure_reason=result.failure_reason,
        )

    return result

這份完整版本保留了 Day 12 和 Day 13 的能力:

  • exact_match
  • contains
  • json_exact

只是會在既有結果上多加:

failure_type

修改 evals/runner.py:寫入 failure_type

接著修改 runner,讓輸出的 eval result 包含 failure_type。

修改 evals/runner.py,找到成功執行 Agent 並 append result 的地方。

原本大概長這樣:

results.append(
    {
        "case_id": test_case["id"],
        "input": test_case["input"],
        "expected": test_case["expected"],
        "grading_method": test_case["grading_method"],
        "task_type": test_case["task_type"],
        "status": "completed",
        "actual": result.answer,
        "passed": evaluation.passed,
        "failure_reason": evaluation.failure_reason,
        "trace_session_id": result.trace.session_id,
        "error": None,
    }
)

修改 evals/runner.py,加入 failure_type:

results.append(
    {
        "case_id": test_case["id"],
        "input": test_case["input"],
        "expected": test_case["expected"],
        "grading_method": test_case["grading_method"],
        "task_type": test_case["task_type"],
        "status": "completed",
        "actual": result.answer,
        "passed": evaluation.passed,
        "failure_type": evaluation.failure_type,
        "failure_reason": evaluation.failure_reason,
        "trace_session_id": result.trace.session_id,
        "error": None,
    }
)

這裡多一個欄位:

"failure_type": evaluation.failure_type

如果測試通過,failure_type 會是:

null

如果測試失敗,則會是:

"format_error"

或:

"wrong_answer"

處理 Agent 執行錯誤

還有一個地方也要修改。

如果 Agent 執行過程發生 exception,程式會進入 except。

例如:

  • API key 沒設定。
  • Gemini API 呼叫失敗。
  • tool call 解析失敗。
  • calculator 收到不合法算式。

這些不是 evaluator 判斷出來的錯誤,而是 Agent 執行流程本身失敗。

所以會分類成:

execution_error

修改 evals/runner.py 的 except 區塊:

except Exception as exc:
    results.append(
        {
            "case_id": test_case["id"],
            "input": test_case["input"],
            "expected": test_case["expected"],
            "grading_method": test_case["grading_method"],
            "task_type": test_case["task_type"],
            "status": "error",
            "actual": None,
            "passed": False,
            "failure_type": "execution_error",
            "failure_reason": "Agent execution error",
            "trace_session_id": None,
            "error": str(exc),
        }
    )

這裡要分清楚 failure_reason 和 error 的差異。

欄位 用途
failure_reason 給 evaluation 統計與報告使用
error 保存實際 exception 訊息,方便除錯

例如:

{
  "status": "error",
  "passed": false,
  "failure_type": "execution_error",
  "failure_reason": "Agent execution error",
  "error": "GEMINI_API_KEY is not set"
}

dashboard 可以把它統計成 execution_error,工程師也還是看得到真正的錯誤訊息。


修改終端機輸出

Day 12 的 runner 已經會印出每題 PASS / FAIL。

可以順手把 failure type 也印出來。

修改 evals/runner.py 的 print_summary():

def print_summary(eval_run: dict[str, Any]) -> None:
    total_cases = eval_run["total_cases"]
    passed_count = sum(1 for result in eval_run["results"] if result["passed"])
    failed_count = total_cases - passed_count

    print(f"Run ID: {eval_run['run_id']}")
    print(f"Total cases: {total_cases}")
    print(f"Passed: {passed_count}")
    print(f"Failed: {failed_count}")
    print()

    for result in eval_run["results"]:
        label = "PASS" if result["passed"] else "FAIL"
        print(f"[{label}] {result['case_id']} - {result['task_type']}")

        if result["failure_type"]:
            print(f"  type: {result['failure_type']}")

        if result["failure_reason"]:
            print(f"  reason: {result['failure_reason']}")

執行後,失敗案例會更容易閱讀。

例如:

[FAIL] case_013 - json_output
  type: format_error
  reason: Output is not valid JSON: Expecting value

執行 Gemini Evaluation

確認你已經設定 Gemini API key:

export GEMINI_API_KEY="你的 Gemini API key"
export GEMINI_MODEL="gemini-3.6-flash"

接著執行:

LLM_PROVIDER=gemini python3 -m evals.runner

這次輸出的結果會多出 failure_type。

例如 JSON 題可能變成:

{
  "case_id": "case_013",
  "input": "請用 JSON 格式回傳 10 + 5 的答案,欄位名稱使用 answer",
  "expected": {
    "answer": 15
  },
  "grading_method": "json_exact",
  "task_type": "json_output",
  "status": "completed",
  "actual": "```json\n{\n  \"answer\": 15\n}\n```",
  "passed": false,
  "failure_type": "format_error",
  "failure_reason": "Output is not valid JSON: Expecting value",
  "trace_session_id": "...",
  "error": null
}

跟 Day 15 相比,這次多了一個可分析的欄位:

"failure_type": "format_error"

現在平台不只知道這題失敗,也知道它是格式錯誤。


以 Day 15 結果做分類

如果用 Day 15 的 Gemini baseline 來看,失敗案例大致可以這樣分類。

case_id task_type 原因 failure_type
case_008 general_qa 回答語意合理,但沒有包含 expected keyword wrong_answer
case_009 general_qa 使用「記錄」而不是 expected 的「紀錄」 wrong_answer
case_013 json_output 外層包了 markdown code fence format_error
case_014 json_output 外層包了 markdown code fence format_error
case_015 json_output 外層包了 markdown code fence format_error

統計後可能會得到:

failure_type 數量
wrong_answer 2
format_error 3

這個結果比單純看:

Failed: 5

更有用。

因為它告訴我們後續改善方向:

  • format_error 可以用 schema instruction、JSON parsing cleanup、retry 或 guardrail 改善。
  • wrong_answer 可以用 prompt 改善,也可以檢查 evaluator 是否太嚴格。

為什麼 case_008 和 case_009 先歸類成 wrong_answer?

Day 15 的 case_008 和 case_009 很值得討論。

例如 case_009:

input: 請用一句話說明什麼是 Trace
expected: 紀錄
actual: Trace 是指記錄並追蹤程式或請求在系統中執行的完整歷程...

這個答案其實有「記錄」,只是 expected 是「紀錄」。

在人類眼中,這很可能是可接受的答案。

但目前 evaluator 是:

contains

它只會做字串包含檢查。

於是「記錄」不等於「紀錄」,這題就會失敗。

這次先把它歸類成 wrong_answer,因為在目前規則下,actual 沒有符合 expected。

但這也提醒我們一件事:

Failure Analysis 不只是在分析 Agent,也是在檢查 evaluator 是否合理。

後面如果要改善,可以有幾種方向:

  • 把 expected 改成多個可接受關鍵字。
  • 加入 synonym normalization。
  • 對開放式回答改用 semantic similarity。
  • 引入 LLM-as-a-Judge。

但這些都不是這一篇的範圍。

先把可統計的 failure_type 欄位建立起來。


檢查 eval run JSON

執行完成後,可以到 data/eval_runs/ 找最新的結果檔。

例如:

data/eval_runs/eval_run_20260906_052327.json

使用以下指令格式化查看:

python3 -m json.tool data/eval_runs/eval_run_20260906_052327.json

你應該會看到每筆 result 多出:

"failure_type": null

或:

"failure_type": "format_error"

通過的案例會是:

{
  "case_id": "case_011",
  "passed": true,
  "failure_type": null,
  "failure_reason": null
}

失敗的案例會是:

{
  "case_id": "case_013",
  "passed": false,
  "failure_type": "format_error",
  "failure_reason": "Output is not valid JSON: Expecting value"
}

重點整理

這次把 evaluation result 從:

{
  "passed": false,
  "failure_reason": "Output is not valid JSON: Expecting value"
}

擴充成:

{
  "passed": false,
  "failure_type": "format_error",
  "failure_reason": "Output is not valid JSON: Expecting value"
}

完成的內容包含:

  • 定義 Agent 常見 failure types。
  • 在 EvaluationResult 加入 failure_type。
  • 新增 classify_failure()。
  • 讓 evaluate() 統一補上 failure type。
  • 讓 evals/runner.py 把 failure type 寫入 JSON。
  • 讓 exception 類型的失敗被標記成 execution_error。

做完後,平台可以回答的不只是:

哪幾題失敗?

而是:

這些失敗分別是哪一類?

這是後續 dashboard、prompt A/B testing、retry 與 guardrails 的基礎。


下一步

Day 17 會把這些 failure type 視覺化。

下一篇會建立錯誤分析表格與第一版 Failure Dashboard,讓我們可以直接看到:

  • 哪些 cases 失敗。
  • 每題的 input、expected、actual。
  • failure type 分布。
  • 不同 task type 的失敗率。
  • 可以點回 trace session 的資訊。

到那時候,failure_type 就不只是 JSON 裡的一個欄位,而會變成分析 Agent 可靠性的第一個 dashboard 指標。


上一篇
Day 15|從 Fake 到 Real:接上 Gemini Flash
下一篇
Day 17|建立錯誤分析表格與 Failure Dashboard
系列文
從黑盒到可驗證:30 天打造 AI Agent 的 Trace、Eval 與 Guardrails 系統 共 17 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言